npj Digital Medicine
○ Springer Science and Business Media LLC
Preprints posted in the last 90 days, ranked by how well they match npj Digital Medicine's content profile, based on 118 papers previously published here. The average preprint has a 0.24% match score for this journal, so anything above that is already an above-average fit.
Knol, L.; Nagpal, A.; Hussain, F.; Beckmann, C. F.; Leow, A.; Eisenlohr-Moul, T. A.; Marquand, A. F.
Show abstract
Digital phenotyping, which is defined as quantifying someone's behaviour with digital devices, provides unprecedented opportunities for understanding human mental health but is hampered by high levels of inter-individual variability. Here, we propose a new method to address this, parsing inter-individual variability by decomposing the digital phenotype dynamics into latent trajectories and using each individual's trajectory membership as a moderator when modelling psychopathology over the same timeframe. We applied our method in the context of mood symptom exacerbation across the menstrual cycle, where symptom severity and timing are inconsistent between individuals. Using the BiAffect platform to collect smartphone typing dynamics, we found stable trajectories in smartphone movement rate: one group of participants showed substantial movement rate fluctuations across the menstrual cycle, whilst the others did not. Participants with movement fluctuations displayed increased fluctuations across the cycle in prospective anhedonia and depression ratings, but not in anxiety, irritability, and suicidal ideation.
Yang, Z.; Zhang, Y.; Love, Z.; Animashaun, A.; Zhong, K.; McDermott, G.; Cai, T.; Liao, K. P.
Show abstract
Objective To develop and evaluate a framework for human-AI interaction. This approach, SHARE (Synergistic Human-Agent REasoning system) was designed to support scalable phenotyping of complex outcomes accurately, robustly and reproducibly from real-world electronic health record (EHR) data to support real-world evidence (RWE) generation. Methods and Analysis Using rheumatoid arthritis (RA) disease activity as the use-case, we studied a multi-institutional EHR-based RA cohort of 3,167 patients. Expert reviewers and a disease activity agent labeled notes using the same review guideline. The agent combined embedding-based informative-note filtering, structured evidence extraction, and evidence-based integrated reasoning to assign disease activity categories with supporting evidence, rationale, confidence, and ambiguity flags. To support scalable deployment, we evaluated a budget-tiered configuration using GPT-5 Nano for high-volume evidence extraction, o4-mini for final reasoning, benchmarking against a GPT-5.4 high reasoning effort configuration applied at every step. Note-level discrepancies were adjudicated by reviewers into final co-produced labels that were used to refine labels and inform agent development. The main outcome measure was the mean absolute error (MAE) of the initial and final agent vs the final co-produced labels. The agreement between agent- and reviewer-flagged ambiguous notes, per-note cost and compute time across configurations were also tested. Results Expert reviewers labeled 626 notes from 273 patients; human-AI adjudication revised 127 (20%) of these initial labels and added 60 newly labeled notes, yielding a 686-note co-produced reference. Against this reference, the final agent's accuracy improved from a mean absolute error of 0.406 to 0.291 with co-learning, and its ambiguity flag agreed with expert ambiguity designations with 92.1% accuracy. Applied across the cohort, the agent labeled 101,691 notes; the budget tiered configuration matched the accuracy of GPT-5.4 at high reasoning effort while reducing estimated cost by 69% and compute time by 70%. Conclusion Adopting a framework for human-AI co-learning, SHARE, improved the overall quality of gold-standard labels, identified ambiguous cases for further review, and supported accurate and standardized chart reviews of disease activity at a scale infeasible for manual review. SHARE's resource efficiency provides a transferable approach to incorporate complex phenotypes in RWE studies.
Firoozbakht, F.; Baumabach, J.
Show abstract
Forecasting a patient's laboratory measurements at future clinical visits from longitudinal electronic health records (EHRs) can support disease monitoring and treatment planning in the context of personalized medicine. However, accurate prediction remains challenging since patients exhibit complex and highly individualized clinical trajectories. Here, we present LaBERT, a transformer-based model trained to forecast future laboratory measurements of a patient given information available at the current and previous clinical visits. Evaluated on 583,535 clinical visits from 255,769 patients in the MIMIC-IV database, LaBERT consistently outperformed baseline methods, reducing mean squared error from 0.77 to 0.53 and improving the coefficient of determination (R2) from 0.29 to 0.51. Medication perturbation analysis further showed that LaBERT learns treatment-related information that is clinically meaningful. In particular, we showed that using the originally prescribed medications, LaBERT predicted future patient states more accurately than when using randomized medication sets in 81% of visits. Furthermore, our controlled counterfactual analyses reproduced established pharmacological effects, including warfarin-associated increases in international normalized ratio (INR) and heparin-associated increases in activated partial thromboplastin time (aPTT), consistently across multiple prediction horizons. These findings establish LaBERT as a model for forecasting future laboratory measurements from longitudinal EHRs and provide a foundation for treatment-dependent patient-state simulation and personalized clinical decision support.
Shi, D.; Shugg, T.; Eadon, M. T.; Su, J.; Chen, Y.; Song, Q.
Show abstract
Objective: Predicting health outcomes from electronic health records (EHRs) is challenging because traditional models rely on structured data and often ignore external medical knowledge. We propose an approach that integrates structured EHR with text-based clinical evidence to improve prediction and interpretability. Methods: We introduce PHO-Agents, a multi-agent system powered by large language models (LLMs) for health outcome prediction. Structured EHR sequences are encoded to produce attention based representations and initial logits, which are converted into patient summaries by a data agent. A retrieval agent gathers relevant clinical guidelines. Research and practical doctor agents independently assess the patient, and a leader agent synthesizes their analyses. Outputs from the EHR based model and the LLM agents are fused to generate final predictions and explanation reports. PHO-Agents was evaluated on three real-world cohorts: acute kidney injury (AKI) patients (in-hospital mortality), chronic kidney disease patients (AKI onset within two years), and cancer patients receiving immune checkpoint inhibitors (immune-related adverse events within one year). Results: PHO-Agents outperformed single-agent and multi-agent LLM baselines across all cohorts. In the AKI mortality task, it achieved a PR-AUC of 90.20 +/- 2.07, compared with 56.46 +/- 2.98 for the best single-agent baseline. Similar gains were observed in the ICI and CKD cohorts. Ablation studies showed that both multi-agent reasoning and logit-level fusion contributed to performance improvements, and case analyses demonstrated clinically consistent explanations. Conclusion: PHO-Agents integrates longitudinal EHR modeling with collaborative LLM reasoning, improving predictive performance, interpretability, and robustness across diverse clinical tasks. This hybrid approach offers a trustworthy strategy for real-world clinical decision support.
Clark, O.; Joshi, K. P.; Joshi, A.
Show abstract
Objective: Online health information seeking is rising, and individuals increasingly act on peer advice without clinical oversight, adjusting doses, delaying care, and modifying treatment. Current misinformation detection assumes factually inaccurate content is what makes these decisions unsafe. We introduce VERITAS (Verification Engine for Risk-aware Information Trust Assessment in health Stories) and formalize the Risk Irrelevance Principle: divergence from accepted clinical practice and potential for harm are distinct, weakly associated dimensions that must be assessed separately. Materials and Methods: VERITAS transforms unstructured health narratives into Agent-Action-Outcome graphs and computes two continuous metrics: Narrative Truth Distance (NTD), quantifying epistemic divergence, and Narrative Risk Score (NRS), assessing harm potential. We evaluated VERITAS on 704 threads from four Reddit health communities. Two domain experts annotated 2,000 segments (Krippendorffs =0.78-0.81). NTD-NRS independence was validated using seven tests. Results: NTD and NRS shared under 5% of variance (r = 0.222; mutual information 0.096 bits): a posts divergence from consensus conveys little about whether acting on it will cause harm. On 435 labeled posts, VERITAS identified 62.2% of expert-labeled misinformation versus 57.5% for the strongest text classifier, the gain concentrated in factually plausible content describing unsafe self-management (27.6% of misinformation) that accuracy-focused classifiers approve. VERITAS assessed 37.8% of this misinformation as low-risk, pending clinical validation. Discussion: Fact-checking-based screening systematically approves the content most likely to prompt unsafe self-management while flagging content least likely to cause harm. Conclusion: Separating divergence from harm potential shifts verification from whether information is correct to whether it is safe to act upon.
Yano, Y.; Shintani, E.; Arita, S.; Ashine, R.; Iinuma, N.; Mori, H.; Fujibayashi, K.; Yamada, Y.; Saita, M.; Nakashima, N.; Itoh, H.; Nangaku, M.; Ohashi, M.; Daida, H.; Arai, H.; Naito, T.
Show abstract
The widespread adoption of clinical large language models (LLMs) introduces significant risks of automation bias, premature closure, and clinician deskilling. Current interpretability paradigms, including latent space trajectories, Concept Activation Vectors, and Concept Bottleneck Models, suffer from topological stagnation, metric distortion, and epistemic occlusion, frequently masking intermediate diagnostic uncertainty behind falsely confident outputs. To address these structural vulnerabilities, this paper introduces a novel closed-loop, multi-agent framework designed to quantify and visualize dynamic epistemic uncertainty in clinical LLM reasoning. By coupling predictive Shannon entropy with non-linear Isometric Feature Mapping (ISOMAP), the architecture projects high-dimensional inference state vectors onto a calibrated two-dimensional latent space, thereby assigning a quantifiable thermodynamic energy state to the reasoning path to track diagnostic velocity, cognitive momentum, and trajectory efficiency across sequential diagnostic rounds. Pilot validation across representative emergency medicine scenarios demonstrated distinct topological and information-theoretic behaviors: unconfounded cases (cerebellar infarction) exhibited smooth geodesic progression toward the ground truth alongside monotonic Shannon entropy decay from 2.15 to 1.74; noisy environments with ambiguous findings (spontaneous pneumothorax) suffered from trajectory wandering, local minimum traps, and high sustained entropy (~2.41) due to insufficient repulsive weighting for negative evidence; and triage-conflicted cases (acute cholangitis) achieved precise geometric proximity to the true node but experienced top-1 rank stagnation because the model conflated acute severity triage (sepsis) with anatomical etiology. By rendering machine hesitation and cognitive divergence visually auditable before final diagnostic crystallization, this geometric-information framework enables dynamic trust calibration and human-AI co-regulation at the point of care while establishing a clear mathematical foundation for future architectural interventions, such as dual-channel safety decoupling and non-linear repulsive weighting. Moving forward, validating these architectural enhancements across large-scale electronic health record databases and prospective clinical trials will be essential to realize its full clinical utility, establishing a foundational blueprint for safe, transparent, and cognitively synergistic AI integration in future medical practice. By rendering the LLM's reasoning process visually auditable, this framework lays the groundwork for capturing and externalizing the clinician's own cognitive patterns within the AI, forming a coupled system. This enables the explicit visualization of cognitive gaps between physician hypotheses and AI inferences, transforming the interaction from simple answer-checking into a dynamic learning process for both human and machine that prevents diagnostic oversight. Ultimately, because the responsibility for final clinical decision-making remains with the human practitioner, this framework serves as a vital decision-support mechanism. Moving forward, validating these architectural enhancements across large-scale electronic health record databases and prospective clinical trials will be essential to realize its full clinical utility, establishing a foundational blueprint for safe, transparent, and cognitively synergistic AI integration in future medical practice.
Desh, S. S.; Achary, P. M.; Nayak, S.
Show abstract
BackgroundSynthetic data generation is increasingly proposed as a strategy to support privacy-preserving data sharing, augmentation of small or restricted biomedical datasets, and benchmarking of artificial intelligence tools in laboratory medicine. However, model selection remains difficult because synthetic data generators differ in fidelity, privacy risk, stability, and generalisability. Existing evaluations have rarely examined performance jointly across conditioning signal strength, synthetic output scale, and train-test generalisation. MethodsWe developed the Synthetic Fidelity-Stability Framework (SFSF), a systematic benchmark of 17 synthetic tabular data generation models using NHANES as a complex biomedical reference dataset. Models included statistical, copula-based, resampling, variational autoencoder, generative adversarial network, and diffusion-based approaches. Synthetic datasets were generated across 11 seed sizes, from 0 to 500 real conditioning observations, and six output scales, from 50 to 5,000 rows, yielding 1,122 synthetic datasets per run. Each dataset was evaluated against the full original dataset, the training subset, and a held-out test subset across five tiers: univariate distributional fidelity, moment agreement, tail behaviour, multivariate dependency structure, and privacy/memorisation risk. Composite rankings and seed-versus-output stability profiles were derived. ResultsUnivariate fidelity was broadly recovered across model classes and was the least discriminating tier. Resampling-based methods ranked highest overall but showed the greatest privacy risk, reflecting proximity to real observations rather than true generative novelty. VAE-family models reproduced moment statistics relatively well but consistently failed on tail and shape fidelity. GAN-family models showed substantial moment-level instability, while VineCopula demonstrated severe multivariate dependency failure. Diffusion-based models, particularly ForestDiffusion, provided the most favourable privacy-utility balance, combining competitive fidelity with the lowest privacy risk and the smallest train-test gap. ConclusionsNo single synthetic data generator dominated across fidelity, stability, and privacy dimensions. The SFSF framework provides a practical, multi-criterion approach for selecting synthetic tabular data generators according to intended clinical laboratory use, balancing statistical realism, dependency preservation, privacy risk, and robustness to seed and output scale.
Rafel, J.; Sartori, D. J.; Finkelstein, H.; Moussa, O.; Triola, M. M.
Show abstract
Problem Authentic patient encounters are the raw material of clinical learning, yet the educational resources learners receive are rarely keyed to the diagnoses in front of them, creating temporal and cognitive gaps. Precision medical education (PME) proposes delivering the right resource to the right learner at the right moment, but practical implementation in the clinical learning environment remains limited. Approach We developed DxMentor, an electronic health record (EHR)-integrated platform that captures each learner's daily inpatient diagnostic exposures from documented International Classification of Diseases, Tenth Revision (ICD-10) codes. Artificial intelligence (AI) is used to match each diagnosis to an educator-curated formulary of micro-learning resources and board-style questions, and to PubMed-derived primary and synthesis literature converted into plain-language evidence summaries. A personalized email "nudge" is delivered before morning rounds, copying supervising attendings for residents, with engagement tracked longitudinally. We report implementation outcomes from July 2024-April 2026. Outcomes DxMentor evaluated 32,846 encounters from 335 medical students and 346 internal medicine residents, delivering 17,340 nudges containing 63,754 didactic resources, 17,038 question sets, and 23,594 summarized articles for approximately $390 in AI token costs. In a benchmarking sample, 91.5% (366/400) of diagnosis-resource pairs were rated relevant by physician-educators. Overall, 78.5% (12,393/15,793) of nudges were opened and 11.3% (1,955/17,340) had at least one click. Engagement was higher among residents than students (open: 80.7% vs. 65.8%; click-through: 12.7% vs. 3.1%; both P < .001), with substantial between-learner variability. Next Steps Email opens and clicks are engagement proxies rather than measures of learning. We are therefore linking nudges to educational outcomes, testing alternative recommendation strategies and timing, and expanding to additional specialties and ambulatory and surgical settings.
Zhang, Z.; Qadir, M. I.; Ramchand, R.; Belwadi, M.; Ball, R. P.; Konstantinopoulos, K.; Abbey, E. M.; Ernsberger, K. T.; Guzman, M. J.; Hendren, S.; Holcomb, B. K.; Robb, B. W.; Stankowski, T.; Waters, J. A.; Stefanidis, D.; Bilimoria, K. Y.; Mohanty, S.; Kolbinger, F. R.
Show abstract
Surgical video interpretation is a promising medical artificial intelligence application. However, no existing video annotation method preserves the spatiotemporal complexity of surgeon reasoning. Here we show that verbal reasoning and visual attention can be converted into structured, machine-actionable records of intraoperative behaviours. Our method decomposes transcribed verbal commentary into video-anchored semantic feedback chunks, which are classified via a large language model, with spatial grounding to surgical scenes via eyegaze or cursor tracking. We demonstrate method validity and scalability on structured and unstructured annotation tasks. For quality feedback on full-length colorectal procedures, the method reached near-human fidelity for chunking (mean cosine similarity: 0.95, SD: 0.01) and semantic classification across observations (mean Cohen's kappa: 0.71, SD: 0.07) and evaluative triggers (mean Cohen's kappa: 0.67, SD: 0.14), with excellent usability ratings. For structured critical view of safety assessment in laparoscopic cholecystectomy, implicit annotation yielded excellent agreement with explicit reviewer ratings (Cohen's kappa: 0.83, 0.49 and 0.81 across three criteria). We anticipate this method will advance surgical data science by enabling scalable construction of meaningfully annotated surgical video datasets.
Bingham, J. C.; Arussy, N.
Show abstract
Active Feature Acquisition (AFA) adaptively selects which diagnostic test to order next and offers a route to reduce unnecessary laboratory testing in acute care. Existing clinical AFA evaluations, however, assume every feature can be retrieved on demand and split data at the visit level, both of which inflate apparent performance. We re-evaluate cost-aware AFA under constraints designed to reflect deployment. From MIMIC-IV we constructed a cohort of 64,766 acute admissions (39,884 patients; 21 conditions; 55 features in 30 test panels) with a patient-level split, a 12-hour decision cutoff, and a per-patient availability mask from what was actually measured, and priced panels using the 2026 Medicare fee schedule under panel-level billing. We evaluated EIG-Cost, which scores each panel by Monte-Carlo Expected Information Gain penalised by its dollar cost, against eight published methods across budgets \30--$60 over five patient-level resamples. At a $30 budget, EIG-Cost achieved the highest macro-F1 (0.188, 95% CI [0.185, 0.191]) at the lowest cost ($17.28), exceeding the strongest baseline in all five resamples (p<0.001; Cohen's d=4.0), and led at every budget. Three of the eight methods collapsed to a vitals-only baseline (macro-F1 approx 0.040), acquiring nothing even at higher budgets, a genuine failure to adapt to availability rather than a budget limitation. Despite modest absolute accuracy, EIG-Cost's probabilities were well-calibrated (expected calibration error $0.048$). Under realistic availability constraints, clinical AFA is substantially harder than full-availability benchmarks imply, several published methods fail outright, and cost-aware information-gain scoring is a robust choice in this harder setting.
Chen, Y.; Zheng, J.; Wang, Y.; Wu, B.; Li, L.; Liu, M.; Xu, L.; Wu, Y.; Liu, C.; Guo, L.; Yang, H.; Bai, X.; Qin, F.; Liao, Q.; Gu, Y.; Zhao, G.; Ma, L.; Pan, K.; Guo, J.; Zhou, Y.; Sun, H.; Tian, Q.
Show abstract
Emergency brain computed tomography (CT) is the first line imaging modality for patients with acute neurological symptoms and trauma, where delayed or incomplete recognition of critical findings can directly compromise clinical outcomes. However, emergency CT interpretation and Chinese reporting remain highly variable under severe time constraints and heterogeneous institutional settings. In this study, we develop ERBrain, a multimodal large model specifically tailored for emergency brain CT, which jointly performs three-dimensional image understanding, Chinese radiology report generation, and emergency severity triage within a unified framework. ERBrain integrates volumetric visual representations with a Chinese large language model and explicitly prioritizes emergency critical signs through risk focused training objectives and a lightweight knowledge-augmented prompting strategy. Using more than 10,000 multicentre emergency CT studies, ERBrain achieved an accuracy of 0.943 and a balanced accuracy of 0.940 for three-level emergency triage and achieved the highest FIES-Avg clinical semantic score among the evaluated report-generation models in the in-distribution cohort. Across external data, ERBrain maintained favourable triage performance in two cross-institutional validation cohorts, whereas performance was lower but remained clinically informative in a third cohort characterized by an extremely low prevalence of Positive cases. These findings support further prospective evaluation of ERBrain as a radiology worklist prioritization and report-drafting assistant in heterogeneous emergency imaging settings.
Tharzeen, A.; Vafaei Sadr, A.; Radfar, N.; Hwang, W.; Abedi, V.; Zand, R.
Show abstract
Background: Machine learning models for stroke mortality prediction typically treat each time horizon independently and use flat tabular features that ignore the relational structure of electronic health records (EHRs). In this pilot study, we leveraged graph-based machine learning models to predict post stroke all-cause-mortality across three different time horizons. Methods: We developed Stroke Temporal Heterogeneous Graph (StrokeTHG), a heterogeneous graph neural network model for simultaneous multi-horizon stroke mortality prediction (30-day, 90-day, 1-year) using EHR data from Penn State Health System. The model encodes various relations among EHR entities (e.g., patient, diagnosis, comorbidity) and temporal encoding of admission time to better predict stroke mortality. We compared our proposed approach against various baseline methods, including Logistic Regression, Random Forest, and XGBoost. We also performed ablation and subgroup analyses, evaluated the quality of learned graph embeddings, and assessed the importance of different edge types in the graph. Results: We included 4,144 stroke patients (mean age 69.2 years; 54.3% men), of whom 3,332 (80.4%) survived their stroke after one year. 30-day, 90-day, and 1-year mortality rates were 9.7%, 13.7%, and 19.6%, respectively. Our proposed approach, StrokeTHG, achieved AUROC of 0.872, 0.878, and 0.837 across horizons, outperforming all tabular baselines. At [≥] , 75% specificity, the model identified 5-10 percentage points more mortality cases than the best baseline at each horizon. Subgroup analysis demonstrated consistent performance across sex subgroups and the largest discriminative gains in the Age 65-80 stratum. Edge-type ablation identified phenotype-patient and admission-patient edges in the constructed EHR graph as the most influential relational edges for mortality prediction. StrokeTHG embeddings outperformed all graph and matrix factorization baselines under an identical downstream classifier, confirming that performance gains stem from representation quality rather than classifier capacity. Conclusions: StrokeTHG demonstrates that heterogeneous graph representations of EHR data provide a consistent improvement over flat tabular models for multi-horizon stroke mortality prediction, with particular advantage at clinically actionable sensitivity thresholds and novel multi-horizon monotonic prediction capability. This methodological framework may be adaptable to other EHR-based clinical research studies seeking to leverage heterogeneous relational structures for predictive modeling.
Yang, J.; Pan, S.; Lim, H. S.; Chu, Y.; Guo, Y.; Agarwal, N.; Babbar, V.; Parikh, G. R.; Chen, Y. T.; Rees, C. A.; Dangor, Z.; Lala, S. G.; Li, Z. R.; Clark, S. J.; Wu, Z.; Datta, A.; Liu, L.; Rudin, C.; Scarpino, S. V.; Gyori, B. M.; McCormick, T. H.
Show abstract
Accurately attributing causes of death is vital for global health, yet fewer than 5% of deaths in resource-constrained regions are medically certified. To assign causes to these unlabeled deaths at scale, practitioners traditionally rely on verbal autopsy, using supervised statistical models to classify based on structured survey data. However, modern mortality surveillance increasingly collects rich, unstructured multimodal data, such as free-text caregiver narratives and postmortem diagnostics, which traditional supervised statistical models struggle to seamlessly integrate. In this paper, we present a comprehensive, multimodal benchmark for cause-of-death classification using data from the Child Health and Mortality Prevention Surveillance (CHAMPS) network, a unique surveillance platform spanning nine countries across South Asia and Sub-Saharan Africa. Using this dataset, we introduce an evaluation framework designed to rigorously assess diagnostic reasoning, moving beyond traditional metrics that fail to capture complex clinical realities. We demonstrate the utility of this benchmark by evaluating zero-shot large language models against supervised baselines across various data modalities. Our results reveal distinct differences in how these modeling approaches synthesize unstructured medical evidence. This benchmark provide a rigorously defined resource for assessing clinical reasoning in next-generation mortality surveillance.
McCann, K. A.; Shin, I.; Li, H.; White, D.; Melnick, E. R.; Iscoe, M. S.; Loza, A. J.
Show abstract
Medical foundation models convert patient records into token sequences for autoregressive prediction, but numeric values such as lab results, vital signs, and time intervals are typically discretized into bins, losing precision and misaligning with clinical thresholds. We trained decoder-only transformer models (47 million parameters) on MIMIC-IV data (364,627 patients; 375 million observations) to compare three tokenization strategies: Discrete (binned values), Continuous Factored (continuous values preserving sequence length), and Continuous Fused (continuous values fused with measurement-type tokens). We evaluated next-token prediction, numeric value prediction, and three clinical tasks: ED disposition at triage, ICD code prediction, and DRG prediction at discharge. Continuous Fused tokenization reduced median sequence length by 34\%, reached the Discrete model's final next-token loss in 30\% of training iterations, and improved numeric prediction accuracy by 30.25\% median nRMSE reduction. ICD code prediction favored Continuous Fused (AU-PRC 0.457 vs.\ 0.446; p < 0.001); DRG prediction was equivalent between Continuous Fused and Discrete; ED disposition accuracy was equivalent across all models ($\sim$0.900), though Discrete achieved better calibration. We additionally explain why predictive performance improves with Monte Carlo sample count and derive a scaling law to predict performance gains from increasing simulation budget. Continuous-value tokenization offers substantial efficiency and precision gains while maintaining comparable clinical task performance, with no modifications to the standard transformer architecture.
Specht, B.; Garbaya, S.; Schneider, R.; Khadraoui, D.; Chavarriaga, R.; Tayeb, Z.
Show abstract
Depression and anxiety are highly prevalent in multiple sclerosis (MS), yet tools for predicting mental health trajectories from clinical data remain limited. We investigated what structured electronic health record data can predict about depression and anxiety progression in MS, and where its limits lie. We developed gradient boosting models to predict PHQ-9 (depression) and GAD-7 (anxiety) score change using EHR data from 2,163 MS patients (7,327 observations) and 1,465 patients (3,319 observations), respectively. Models achieved R^2 of 0.22 (PHQ-9) and 0.28 (GAD-7). Baseline score was the dominant predictor, but this largely reflects regression to the mean: patients with high baseline scores tend to improve, while those with low scores tend to worsen. Age emerged as a consistent secondary predictor across both models: younger patients showed smaller improvements independent of baseline severity. Feature importance differed between models---PHQ-9 prediction relied on symptom subscales while GAD-7 incorporated pain and disease duration. These results suggest that structured clinical data alone capture only a fraction of what drives mental health trajectories, and that richer data sources---clinical notes, patient-reported outcomes, digital phenotyping---will be needed to enable meaningful individual-level prediction.
Radoynova, M.; Benouis, M.; schulze, f.; Winter, S.; Bornhauser, M.; Middeke, J. M.; Eckardt, J.-N.
Show abstract
Large Language Models (LLMs) are increasingly used by clinicians and patients for medical queries, yet their accuracy and safety at the specialist level in hematology remain insufficiently characterised. We benchmarked ten frontier proprietary and open-weight LLMs across two generations on 1,477 board-style hematology multiple-choice questions (MCQs) derived from five educational datasets spanning nine disease areas and six clinical skill domains, including text-only and multimodal case vignettes. Claude Opus 5 had the highest mean accuracy (92.7% text, 76.9% multimodal), followed closely by Gemini-3.1 Pro (91.4% and 78.7%), Gemini-3.6 Flash (91.0% and 74.8%) and GPT-5.6 Sol (89.9% and 76.7%). Accuracy significantly correlated with model size both for text-only and multimodal MCQs. Between model generations, the largest improvements in accuracy were seen for open-weight models whereas proprietary models showed only marginal gains. In error analysis, top-performing models exhibited highly concordant failure patterns, suggesting shared limitations on challenging cases. Frontier LLMs exhibit substantial specialist hematology knowledge across diverse subspecialist domains and clinical skill sets. Yet, despite high accuracy on board-style questions in hematology, continuous expert-on-the-loop output monitoring is paramount.
Proulx, J.; Daines, B.; Barton, M.; Leonard, M. E.; Garcia, J. A.; Young, B.; Snell, Q.; West, T. W.; Watson, S. R.; AlQaseer, M.; Louiset, M.; Maqsood, M. B.; Voutt-Goos, M. J.; Douma, C.; Kasbekar, N.; Jeffries, J.; Abu-Rahmeh, W.; Frush, K.; Grewal, D. K.; Bahsoun, M.; Leonard, M.; Frankel, A.; Classen, D. C.; Pestotnik, S. L.
Show abstract
Objective. To introduce PsiBench, a clinically validated medication-safety benchmark for evaluating large language models (LLMs) against the standards used to certify hospital computerized provider order entry (CPOE) and electronic health record (EHR) systems, and a non-overlapping three-tier evaluation framework separating highest-stakes discrimination, the operational CDS regime, and category-correct alerting. Materials and Methods. PsiBench comprises 492 medication-safety scenarios across 11 safety categories, created by clinical pharmacology experts whose work underpins an annualized testing procedure used by more than 2,000 U.S. hospitals. The three-tier framework partitions the scenarios non-overlappingly: Discrimination (98 scenarios, 50 fatal vs 48 deception, near-balanced 51%/49%); Operational (394 scenarios, 261 serious unsafe plus 133 safe including 41 Excessive Alerts reclassified as operational negatives); and Attribution (311 alert-required scenarios). We evaluated 40 frontier LLMs from 10 providers over 3 runs per scenario at temperature 0.2 (or the provider default where temperature is not configurable), yielding 59,040 evaluations conducted April 21-23, 2026. Results. Headline binary performance on the full benchmark spans a wide range across the 40 models: F1 78.5%-92.3%, accuracy 65.4%-89.8%, sensitivity 81.4%-100.0%, specificity 6.1%-81.8%. Leading models by F1 (o4-mini 92.3%; o3 92.2%) pair high sensitivity with meaningful specificity; three models saturate sensitivity at 100% but fall below 25% specificity, indistinguishable from a naive always-alert classifier. The wide spread on a single headline metric motivates tier-specific analyses, developed in a separate clinical paper. Discussion and Conclusion. PsiBench and the three-tier framework operationalize a rigorous evaluation rubric for LLM medication safety, grounded in two decades of national hospital audit experience. The framework generalizes to any binary medication-safety classifier (rule-based, conventional ML, or LLM-driven), supporting tier-aware model selection and post-deployment surveillance.
Levy, J.; Levis, M.; Dimambro, M.; Rozema, L.; Ayandeh, S.; Diallo, A.; Zhou, Y.; Li, S.; Wu, W.; Shiner, B.; Gui, J.
Show abstract
Background: Suicide remains a significant and potentially preventable cause of death among United States veterans. Predictive models based on structured electronic health record (EHR) data, including the U.S. Department of Veterans Affairs' Recovery Engagement and Coordination for Health-Veterans Enhanced Treatment (REACH-VET) program, aim to identify individuals at elevated risk for enhanced monitoring and follow-up. Increasing evidence suggests that unstructured clinical narratives contain additional psychosocial information that may enhance risk prediction when analyzed using natural language processing (NLP). However, optimal approaches for representing clinical text remain uncertain. Recent advances in large language models (LLMs) enable contextual text representations that capture complex semantic relationships beyond traditional lexical methods. Methods: We compared the predictive performance of pretrained LLMs with classical bag-of-words (BoW) representations for suicide risk prediction using clinical notes from 27,241 veterans receiving care in the Veterans Health Administration. Patients were stratified by REACH-VET risk tier (low, moderate, high), and models were evaluated across prediction windows defined by note look-back periods (<30, <90, and <270 days). Results: LLM-based representations outperformed BoW approaches in seven of nine risk tier-time window combinations, achieving a maximum AUROC of 0.644 when solely considering text. Incorporating structured clinical variables further improved performance (AUROC=0.748). Model interpretation identified suicide-related language, especially in notes documented within 30 days of the outcome among patients classified as high risk. Conclusions: Pretrained LLMs can extract clinically meaningful information from narrative documentation, providing a foundation for future work adapting to additional clinical contexts and nuanced temporal associations to improve suicide risk prediction.
Hayder, N. S.; Bukhari, S. A. C.
Show abstract
Synthetic clinical data are increasingly used for healthcare machine-learning development, model validation, data sharing, and predeployment testing, yet such data often claim to be trustworthy after passing a limited collection of realism tests. A synthetic dataset may indeed claim statistical similarity while leaking training membership, erasing rare subgroups, failing on held-out real patients, or lacking sufficient artifacts for reproduction. We introduce SynTrustBench, an evidence-gated and executable benchmark for evaluating trustworthiness claims across five non-compensable dimensions: fidelity, clinical utility/validity, privacy, equity, and robustness/generalization. Its Evidence Assessment component audits published reports and produces a five-element Evidence Maturity Profile (EMP) together with a separate evaluability gate. Its executable structured-tabular protocol accepts frozen real training data, held-out real test data, a synthetic table, and a declarative configuration; computes dimension-specific metrics and uncertainty; and produces subgroup results, failure flags, benchmark cards, and provenance manifests. In a frozen pilot audit of 30 reports, 17 of 30 quantitatively evaluated privacy, 2 of 30 documented a formal privacy guarantee to the audit threshold, 2 of 30 evaluated equity, 12 of 30 evaluated robustness, and only 4 of 30 passed the evaluability gate. The executable implementation operationalizes the same dimensions through distribution and dependency checks, frozen train-on-real/test-on-real (TRTR) and train-on-synthetic/test-on-real (TSTR) utility, empirical privacy attacks, subgroup analysis, perturbation testing, and a controlled failure-injection harness. SynTrustBench does not certify clinical safety or collapse trustworthiness into a single score. Instead, it provides an inspectable predeployment contract for identifying what was evaluated, what failed, what remains unknown, and whether evidence is sufficiently complete and reproducible for comparison or downstream healthcare AI use.
Tao, J.; Fenn, N.; Parent, H.; Wu, H.; Arnold, T.; Etue, J.; Chen, E.; Chan, P.
Show abstract
Background: Depression and anxiety are managed largely between clinical visits, yet outpatient care lacks scalable, accountable mechanisms for between-visit support. Large language models converse fluently but fuse clinical reasoning with language generation in one opaque process, so they cannot reliably deliver evidence-based psychotherapy and typically operate outside clinician oversight. Objective: To evaluate C-Mind, a provider-supervised neuro-symbolic system in which a Clinical Knowledge Graph (KG) governs therapeutic decisions for a large language model across eight psychotherapy modalities. Methods: Two simulation regimes addressed eight pre-specified governance questions: a structural validation of KG routing against 117 guideline-anchored vignettes, and a governance battery using progressively disclosing LLM patient agents to evaluate decision traceability, repeatability, provenance auditability, adversarial crisis-detection robustness (277 probes), provider treatment-goal governance, and counselor technique adherence. Crisis detection was additionally validated externally against an independent, clinician-annotated corpus (CRADLE Bench). Results: The KG routed 116/117 vignettes (99.1%) to guideline-appropriate care and detected all 18 high-risk presentations, firing a therapy-suppressing hard halt on 16/18. Adversarial crisis-detection sensitivity was 96.7% and specificity 95.4% (277 probes); on external validation, the system detected 98.5% of 600 dialogues with ongoing suicidal ideation or self-harm at or before the annotator confirming turn. Decisions were 99.1% repeatable, 100% reconstructable per turn, and 100% provenance-auditable across all 354 KG nodes. Provider-set diagnosis, goals, and safety context governed behavior deterministically. Stripped of governance, the same model produced unsolicited clinical monologues on 100% of turns (vs 9% governed) and delivered diagnoses and medication advice the governed system never produced. Conclusions: A neuro-symbolic architecture achieves near-perfect guideline-appropriate routing with a governance profile, traceability, reproducibility, machine-traceable provenance, externally validated crisis detection, and deterministic provider control aligned with requirements for regulated clinical AI.